Papers with iterative offline RL training

1 papers
Preference-Guided Reflective Sampling for Aligning Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Repeated random sampling is a widely used method that independently queries the model multiple times to generate outputs.
Approach: They propose a more efficient method for iterative data generation and model re-training that leverages tree-based tree-derived generation framework to enable more efficient sampling.
Outcome: The proposed method significantly outperforms repeated random sampling in best-of-N sampling on AlpacaEval and Arena-Hard.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations